Papers with LLM-based chat-bots
ChatBench: From Static Benchmarks to Human-AI Evaluation (2025.acl-long)
Copied to clipboard
| Challenge: | In 2024, 40% of US adults reported using generative AI in their everyday lives, an unprecedented rate of adoption for a new technology. |
| Approach: | They propose to convert MMLU questions into user-AI conversations by seeding the user with the question and having them carry out a conversation with the LLM to answer their question. |
| Outcome: | The proposed model can estimate user-AI accuracy by fine-tuning a user simulator on a subset of ChatBench. |